Papers with human-robot interaction

8 papers
Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction (L18-1)

Copied to clipboard

Challenge: a study investigates the influence of gender stereotypes on trust and likability of humanoid robots . explicit gender and stereotypicality of a task are manipulated to influence robot behavior . future research may look into situational variables that drive stereotypification in robot interaction .
Approach: They investigated the influence of gender stereotypes on trust and likability of robots . they used explicit (name and voice) and implicit (personality) genders to manipulate stereotypical tasks . future research may look into situational variables that drive stereotypization .
Outcome: The findings suggest that gender stereotypes need to be differentiated in robot interaction . the gender and personality characteristics of robots influence trust and likability .
Dialogue-AMR: Abstract Meaning Representation for Dialogue (2020.lrec-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context.
Approach: They propose a schema that enriches Abstract Meaning Representation (AMR) it provides a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems.
Outcome: The proposed schema provides a semantic representation for facilitating Natural Language Understanding (NLU) in human-robot dialogue systems.
Learning Physical Common Sense as Knowledge Graph Completion via BERT Data Augmentation and Constrained Tucker Factorization (2020.emnlp-main)

Copied to clipboard

Challenge: Physical commonsense learning is an essential part of human-robot interaction . existing methods of learning physical commons sense suffer from generalization .
Approach: They propose to use physical commonsense learning as a knowledge graph completion problem to better use latent relationships among training samples.
Outcome: The proposed method outperforms existing methods in the human-robot interaction problem.
Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Currently, image retrieval systems can retrieve relevant results for diverse inputs, but they do not provide a way to intentionally inject variety into the search results.
Approach: They propose a multimodal dataset that combines semantic annotations with image bounding boxes.
Outcome: The proposed system improves image retrieval performance and flexibility.
FARMI: A FrAmework for Recording Multi-Modal Interactions (L18-1)

Copied to clipboard

Challenge: a new framework for recording multi-modal data is needed to capture multi-party, richly recorded corpora and perform real-time processing of such data.
Approach: They propose an open-source processing architecture for corpora and real-time processing . they deploy the architecture in a multi-party deception game with six humans and one robot .
Outcome: The proposed architecture is agnostic to hardware and programming languages, although it's mostly written in Python.
GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented Collaborations (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on the use of exocentric and egocentric videos in video question answering are focusing on eye-gaze information.
Approach: They propose a task-oriented VQA dataset that captures eye-gaze information . they propose assisting models that ground the perceptual input into semantic information based on three different answer types .
Outcome: The proposed model can ground the perceptual input into semantic information while reducing ambiguities.
How Much Do Robots Understand Rudeness? Challenges in Human-Robot Interaction (2024.lrec-main)

Copied to clipboard

Challenge: This paper examines the pressing need to understand and manage inappropriate language within the evolving human-robot interaction landscape.
Approach: They propose to use data cleaning methods to identify inappropriate language in real-time interactions and evaluate natural language models for their proficiency in discerning rudeness.
Outcome: The proposed methods identify and mitigate inappropriate language in real-time interactions and evaluate natural language models for their proficiency in discerning rudeness.
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory (2025.acl-long)

Copied to clipboard

Challenge: Existing ground VLN agents struggle in aerial VLLN due to the lack of predefined navigation graphs and the exponentially expanding action space in long-horizon exploration.
Approach: They propose a large language model-empowered aerial VLN agent that decomposes the long-horizon task into sub-goals with different semantic levels.
Outcome: The proposed method achieves state-of-the-art performance with significant improvement in continuous city environments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations